Papers with Gemini 2.5
To Mask or to Mirror: Human-AI Alignment in Collective Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used to model and augment collective decision-making. |
| Approach: | They propose a framework for assessing collective alignment using the Lost at Sea social psychology task. |
| Outcome: | The proposed framework compares LLMs with human-AI alignment on the Lost at Sea social psychology task. |
StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work focused on typographic and pixel-level perturbations, leaving the study of SCO unexplored. |
| Approach: | They propose a framework that exploits MLLMs' diagrammatic reasoning capabilities to bypass safety guardrails. |
| Outcome: | The proposed framework exploits the model's reasoning capabilities to bypass safety guardrails. |
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts (2025.emnlp-main)
Copied to clipboard
| Challenge: | Scientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts. |
| Approach: | They propose a benchmark for scientific fact-checking grounded in scientific charts . climateViz comprises 49,862 claims paired with 2,896 visualizations . results show current models struggle to perform fact- checking when statistical reasoning is required . |
| Outcome: | The climateviz benchmark is the first large-scale benchmark for scientific fact-checking . it includes 49,862 claims paired with 2,896 visualizations labeled as support, refute, or not enough . |